Medical Physics
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match Medical Physics's content profile, based on 14 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Chen, W.-Y.; Wan, S.-Y.; Lin, G.-Y.
Show abstract
Accurate segmentation of thin-wall organs-at-risk (OARs)-the cochlea, vestibular semicircular canals, internal auditory canal, tympanic cavity, and middle ear-is clinically relevant for head-and-neck radiotherapy planning, yet these small, thin-wall structures remain among the most challenging targets for automated delineation. Dual-frequency feature fusion is a promising direction for boundary-sensitive representation, but under the investigated FP16 FFT-FcaNet setting, we observe an approximately 863-fold activation-scale mismatch between the FFT and FcaNet branches, causing a nominal 5 percent residual coefficient to behave as an approximately 43-fold dominant term. We propose FreqFuseNet, which resolves this mismatch by normalizing the FcaNet branch to the FFT activation scale before residual injection with a fixed low-amplitude coefficient (beta = 0.05), restoring beta as an interpretable 5 percent residual-amplitude coefficient relative to the FFT feature scale. Under a controlled binary per-OAR ROI protocol on the SegRap2023 head-and-neck CT benchmark across 10 clinically prioritized thin-wall OARs, FreqFuseNet achieves Dice of 0.849, HD95 of 0.824 mm, and SDice@1mm of 0.959 in the primary seed, with comparable performance in an independent second seed (Dice 0.843, HD95 0.823 mm). FreqFuseNet yields statistically significant case-level aggregate improvements over 3D U-Net and MedNeXt-S (Wilcoxon p < 0.01 and p < 0.05, respectively), using only 29.7 million parameters versus 414.6 million for the full wavelet baseline.
Hamkins, H. M.; Tam, K. H.; Sobremonte, A.; Jogi, S.; Koay, E.; Hassanzadeh, C.; Segars, P.; Tyagi, N.; Subashi, E.
Show abstract
Background: Independent end-to-end verification of adaptive radiotherapy on MR-Linac systems is limited by the lack of patient-specific phantoms able to reproduce imaging and dosimetric properties from CT and MRI scanners. We present a method for automated generation of 4D, patient-specific, multi-material 3D-printable phantoms for quality assurance of adaptive radiotherapy on a 1.5T MR-Linac. Methods: Patient images were automatically segmented using a pretrained deep learning model. The segmented structures were converted into high-resolution 3D meshes and assembled into printable phantoms. A dosimeter holder was inserted at user-defined anatomical locations, with orientation optimized to avoid traversal across heterogeneous tissue interfaces. Physiological motion was incorporated by generating phantoms from images at different timepoints and interpolating deformation fields to create continuous 4D models. Multi-material organs designed by mixing a set of six polymers at various proportions were used to reproduce tissue-specific imaging properties. The properties of material mixtures were evaluated in a clinical CT simulator and a 1.5T MR-Linac. Results: The proposed workflow enables automated generation of anatomically realistic phantoms with several types of embedded dosimeters. A discrete search method was designed for placement and immobilization of OSLD, film, and ion chamber dosimeters. Calibration curves for Hounsfield units were derived through variations in radiopaque material content, while MR signal intensity was modulated by gel and tissue matrix mixtures. Patient-derived abdominal phantoms were fabricated at multiple scales while replicating internal anatomical detail. Multi-dimensional phantom generation enabled continuous representation of motion states with consistent mesh topology across phases. Conclusions: We demonstrate an end-to-end workflow for automated generation of 4D patient-specific phantoms for MR-Linac quality assurance. The method combines realistic anatomy, embedded dosimetry, multimodal imaging properties, and physiological motion within a single fabrication framework. This approachmay enable an improved validation of adaptive radiotherapy workflows in MR-guided treatment devices.
Pasyar, P.; Mei, K.; Im, J. Y.; Roshkovan, L.; Geagan, M.; Noël, P. B.
Show abstract
ABSTRACT Background: Metallic implants such as orthopedic screws, prostheses, and dental hardware produce beam-hardening, photon-starvation, and streak artifacts that degrade computed tomography (CT) image quality, and the metal artifact reduction (MAR) methods developed to mitigate them require objective, reproducible benchmarking. Purpose: Objective evaluation of MAR algorithms in CT is hindered by the absence of phantoms that simultaneously provide anatomically realistic backgrounds, embedded implants of known geometry, and controllable, ground-truth--referenced artifact intensity. We present a dual-filament, voxel-level three-dimensional (3D) printing method that fulfills these requirements and demonstrate its capabilities on a clinically representative cervical spine case with embedded orthopedic spinal screws. Methods: The proposed method extends the PixelPrint framework, a fused-deposition-modeling (FDM) workflow that converts clinical Digital Imaging and Communications in Medicine (DICOM) data directly into 3D-printer Geometric code (G-code) without intermediate segmentation or surface meshing, to interleaved, voxel-level deposition of two filaments: a calcium-doped polylactic acid (PLA) for soft tissue and bone, and a higher-attenuation metal-doped PLA for metallic implants. For demonstration, anonymized DICOM data of a healthy cervical spine were used to design and fabricate three matched phantoms, each with six embedded spinal screws at C4--C6: a 0% metal-infill ground-truth phantom, a 50% medium-metal-infill phantom, and an 85% high-metal-infill phantom. All phantoms were scanned on a clinical spectral CT system at 120 kVp and 1000 mAs, reconstructed at 0.67 mm slice thickness with virtual monoenergetic imaging (VMI) across 50--190 keV. Method performance was characterized by region of interest (ROI)-based Hounsfield Unit (HU) agreement with the source patient data and by the noise-independent Gumbel-distribution p-index metric. Results: The dual-filament method reproduced patient anatomy, soft-tissue contrast, and screw geometry with high fidelity. ROI HU values agreed with patient data within {+/-}25 HU for soft tissue and trabecular bone; cortical regions were underestimated owing to the current ceiling of the calcium-doped PLA used in this study. The tunable-artifact behavior was quantified as follows: the Gumbel location parameter scaled monotonically from 46.7 HU (no-metal background) to 57.1 HU (50% infill) to 90.5 HU (85% infill) for the VMI 70 keV with standard filter. High-keV VMI reconstructions substantially reduced streak and beam-hardening artifacts while preserving anatomic detail. Conclusions: The proposed dual-filament, voxel-level PixelPrint method enables the fabrication of patient-specific, multi-material CT phantoms with embedded metallic implants and controllable, ground-truth--referenced artifact intensity. Although demonstrated here in a single cervical-spine case, the workflow is anatomy- and implant-agnostic by construction and could in principle be adapted to other musculoskeletal sites (e.g., knee, hip, dental) and implant materials, providing a reproducible methodological foundation for benchmarking MAR algorithms, characterizing spectral CT performance, and validating emerging photon-counting detector systems. Keywords: 3D printing methodology; fused deposition modeling; voxel-level multi-material printing; spectral computed tomography; metal artifact reduction; phantom design; orthopedic implants; dual filament; PixelPrint.
Genske, U.; Laudani, A.; Yan, L.; Peng, Y.; Boening, G.; Ulas, S. T.; Wagner, M. P.; Diekhoff, T.; Hamm, B.; Jahnke, P.
Show abstract
Artificial intelligence (AI) applications in computed tomography (CT) imaging require objective and continuous testing, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised testing and monitoring of AI, demonstrated in liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that anatomically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.
Yamamoto, S.
Show abstract
CT perfusion (CTP) is central to acute-stroke and oncologic imaging, yet quantitative outputs vary substantially across vendor software, undermining reproducibility. We present an open, transparent core (ctp-core) that fits first-pass time-attenuation curves with a gamma-variate model, derives perfusion indices (peak enhancement, time-to-peak, bolus-arrival time, and area under the curve) analytically from the fitted parameters, and renders parametric maps with the ASIST-Japan standardized lookup table (a-LUT) so that visualization is comparable across sites. Every parameter, bound, and processing step is exposed. The method is validated on Monte-Carlo synthetic curves with known ground truth; no confidential or patient data are used. Across signal-to-noise ratio (SNR) levels 5 to 100 (200 independent runs per level) the pipeline recovers peak time to within 0.03-0.52 s and peak amplitude to within 0.4-8.1% (mean absolute error), degrading monotonically with noise; at a representative SNR of 20 it recovers peak time within 0.13 s, peak amplitude within 2.0%, and bolus-arrival time within 0.51 s, with fit quality R-squared = 0.98. The reproducibility demonstration is deterministic (fixed seed) and re-runs to bit-stable metrics. All code, the synthetic-data generator, the standardized-visualization module, evaluation scripts, and a 34-test suite are released openly for independent verification. The contribution is a fully open, parameter-transparent gamma-variate plus standardized-visualization pipeline with reproducible synthetic benchmarks: a reference others can audit, reuse, and build on.
Gu, X.; Zhu, H.; Zhong, F.; Teng, G.-J.
Show abstract
Background: Nuclear medicine and radiopharmaceutical development require coordinated radiochemistry, dosimetry, molecular imaging, radiation-safety and clinical decision processes. Current workflows remain fragmented, difficult to audit and poorly standardised for evaluating domain-specific AI support. Methods: We developed RadGuide AI, a nuclear medicine agent built around a traceable data-model-tool loop. Patent, literature and clinical-trial records were converted into 15,596 initial QA items; relevance screening, completeness checks, semantic deduplication and cross-validation retained 5,474 core QA items. MedGemma-27B-Instruct served as the foundation model and was adapted with LoRA. The system incorporated 55 MCP-wrapped tools covering radiopharmaceutical R&D, clinical decision support, imaging analysis and radiation-safety/dosimetry. Evaluation used a locked N=200 benchmark with predefined denominators, leakage control, expert scoring, statistical procedures, factuality audits and tool-execution metrics. Results: RadGuide-LLM achieved 88.5% answer accuracy (177/200; 95% CI, 83.3-92.2%) and a Macro-Average score of 21.5/25 (bootstrap 95% CI, 20.9-22.0), exceeding GPT-4o, DeepSeek-V3.2 and the base MedGemma model in this technical evaluation. Supplementary audits reported guideline compliance, terminology recall, knowledge coverage, tool-routing success and preclinical/phantom dosimetry agreement with explicit denominators and confidence intervals. Interpretation: RadGuide AI converts nuclear medicine queries into auditable retrieval, tool selection, calculation, verification and reporting workflows. The findings support technical feasibility, not definitive patient-level clinical validation; prospective multicentre studies and external benchmark release remain required before clinical deployment.
Rich, J. M.; Kang, R.; Jin, D.; Subramanian, S.; Duddalwar, V.; Pachter, L.
Show abstract
We developed a standardized, reproducible preprocessing framework for computed tomography (CT) imaging data from multi-institutional repositories such The Cancer Imaging Archive (TCIA), enabling consistent radiomics and artificial intelligence (AI) analyses. Imaging data from TCGA-KIRC patients available on TCIA were used as a representative heterogeneous dataset characterized by variation in acquisition protocols, inconsistent metadata, and differing image quality. The proposed modular pipeline includes series filtering, DICOM-to-NIfTI conversion, orientation harmonization to a canonical coordinate system, voxel spacing normalization, intensity clipping and normalization, segmentation integration, and metadata validation, and is implemented in a reproducible, notebook-based framework compatible with common radiomics and deep learning workflows. This pipeline standardizes imaging data into analysis-ready volumes with consistent geometry, intensity distributions, and spatial alignment, reducing non-biological variability that can adversely affect radiomic feature stability and model performance. The modular design enables task-specific adaptation of individual preprocessing steps while maintaining overall consistency. Although demonstrated on TCIA, this framework is generalizable to other heterogeneous imaging datasets and provides a foundation for robust, large-scale computational imaging studies.
Knol, M.; Goncalves Jorge, P.; Kunz, L. V.; Korysko, P.; Petit, B.; Durham, A.; Marie-catherine, V.; Tsoutsou, P.; Koutsouvelis, N.; Lascaud, J.
Show abstract
Objective: Preclinical small-animal irradiators such as the FLASH-SARRP can support the advancement of photon-FLASH toward the clinic. This study aimed at characterizing the FLASH-SARRP and established a robust quality assurance (QA) workflow to enable accurate and reproducible preclinical experiments. Approach: Custom 3D-printed spacers were designed to ensure reproducible X-ray tube alignment, sample positioning and mounting of the dosimetric tools. Beam characteristics were evaluated using a combined dosimetric approach. High spatially resolved dose distributions were obtained from Gafchromic films, whereas a plastic scintillating fiber was employed to monitor in real-time the temporal pulse structure and synchronization between the two X-ray tubes. Day-to-day variability of the delivery was evaluated over several sessions. Main results: The FLASH-SARRP achieved dose-rates of around 80 Gy/s when both tubes were used simultaneously and provided a homogeneous irradiation field suitable for small-animal studies. A desynchronization between the two tubes was observed with an average delay of 10 ms, resulting in temporal dose-rate heterogeneity. Additionally, a substantial inter-session variability (~11%) was found, whereas the intra-session variability was relatively low (~4%). Inter-session variability was reduced to 5%, approaching the intra-session variability, by adding Gafchromic films/scintillator-based quality assurance (QA) workflow into the irradiation routine. Significance: This work highlights the importance of temporal dosimetry for preclinical FLASH studies. Additionally, a practical QA framework is proposed integrating real-time monitoring with reference dosimetry. The proposed work enables adaptive dose delivery, thereby enhancing the reproducibility of the irradiations, which is crucial for reliable preclinical studies on the FLASH effect.
Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.
Show abstract
Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.
Arndt, M. D.; Hansler, R.; Tirinato, L.; Tkachenko, A.; Seco, J.; Schepers, U.; Spadea, M. F.
Show abstract
Background: Three-dimensional tumor spheroids are an established radiobiology model, but scalable, reproducible readouts of dose-dependent radiation response are lacking. We evaluated whether optical coherence tomography (OCT) radiomics can quantify dose-associated response in spheroids, and how it compares with conventional brightfield morphology. Methods: This in vitro, cross-sectional study used SAS oral squamous cell carcinoma spheroids seeded at two densities (5000 and 10000 cells), irradiated at 0 to 12 Gy, and imaged on days 1 to 11 post-irradiation. Each OCT acquisition yielded co-registered structural-intensity and speckle-variance volumes. Radiomic features (shape, first-order, texture) were extracted with Radiomics.jl, filtered for repeatability, correlation-pruned, and ensemble-ranked. Dose correlation was assessed by repeated 5-fold cross-validation across five regressors, comparing brightfield-only (BF), OCT-only, and combined OCT+BF feature sets with paired Wilcoxon tests. Results: OCT-only models consistently outperformed the BF baseline (median R2 0.77 to 0.85 versus 0.61 to 0.69; p<0.001 for all regressors). Adding brightfield to OCT gave no consistent benefit, reaching significance only for Random Forest (p=0.026, power 0.62). A compact shared feature subset combined brightfield area dynamics with OCT texture, shape, and speckle-variance descriptors, all showing low repeat-scan variability relative to cohort variability. Conclusions: OCT radiomics provides a sensitive, reproducible, label-free high-throughput readout of spheroid radiation dose response that outperforms the current brightfield-based approach, without requiring concurrent brightfield acquisition.
Shanbhag, A.; Miller, R. J.; Killekar, A.; Marcinkiewicz, A. M.; Zhou, J.; Lemley, M.; Kamagate, A.; Van Kriekinge, S. D.; Kavanagh, P. B.; Feher, A.; Miller, E. J.; Liang, J. X.; Berman, D. S.; Dey, D.; Leahy, R. M.; Slomka, P.
Show abstract
Background: Coronary artery calcium (CAC) is an established measure of coronary atherosclerosis from computed tomography (CT). While deep learning (DL) can quantify CAC from non-dedicated CT, the accuracy is limited by image quality. Purpose: We derived and validated a novel method for DL CAC segmentation on ultra-low dose CT attenuation correction (CTAC) scans that is trained with synthetic low-dose, ungated images. Materials and Methods: Models were trained using one center and externally tested in two other centers. Synthetic, ungated CT scans were generated so that expert segmentations from dedicated CAC scans could be used as ground truth for perfectly registered synthetic images through knowledge adaptation (KAD-CAC). We evaluated agreement between CAC scoring methods vs expert readers on a per-patient and per-vessel basis, as well as associations with the primary outcome of death or myocardial infarction (MI). Results: The DL models were externally tested on 5969 patients with a median age of 64 (IQR 56 - 73), of whom 50.2% were male. The KAD-CAC model had higher Cohens kappa K (0.86, 95% CI 0.85 - 0.87) compared to previous convolutional LSTM model (K 0.78, 95% CI 0.76 - 0.80, p<0.01), or models trained with only gated images (K 0.81, 95% CI 0.80 - 0.82, p<0.01). Net reclassification improvement for CAC stratified risk of death or MI, was greatest for the KAD-CAC model over baseline including age, sex, hypertension, diabetes, dyslipidemia, family history, smoking, stress total perfusion deficit, and left ventricular ejection fraction. Conclusion: We use paired synthetic ungated scans to transfer expert gated CAC annotations into the ungated domain, resulting in substantially better vessel-level CAC scoring and improved risk stratification.
Wang, J.; Tang, W.; Ma, X.; Yan, H. m.; Yuan, Y.
Show abstract
Large language models (LLMs) are increasingly used for automated quality control (QC) of radiology reports. However, the reliability of LLMs on reports in Mandarin, and the relative performance of domestic versus international flagship models, remain unknown. We benchmarked 14 LLM configurations, seven Chinese-developed ("domestic") and seven international models, on 1,000 whole-body 18F-FDG PET/CT reports split into an error-injected "junior-docto" arm and a low-residual "finalised" arm (500 each), using a controlled error-injection gold standard. Under each blinded zero-shot prompt, each model flagged six error types and assigned a 1-5 overall score. Two distinct abilities: error-detection macro-F1 (0.356-0.667) and overall-score calibration (ICC[2,1] 0.099-0.627), were weakly and not significantly correlated across models (Spearman {rho} = 0.38, p = 0.18); the dissociation was instead evident in sharp rank reversals, the strongest detector (Claude-Opus-4.8 0.667) calibrating poorly (0.491), while the three best-calibrated models were all domestic (MiMo 0.627, GLM-5 0.612, DeepSeek 0.609). Once the access channel was controlled, domestic and international error detection were statistically indistinguishable ({Delta}macro-F1= -0.011, P = 0.84); domestic models showed consistent but not significant advantages in calibration ({Delta}ICC = +0.142) and Chinese-character-error detection ({Delta}F1 = +0.109), accompanied with large reductions in cost (US$0.09-2.71 vs $0.26-14.5 per 1,000 reports) and on-premise deployability. Re-running two flagships through both agent channels and clean APIs showed that agent channel inflated both detection and calibration (GPT-5.5 {Delta}ICC = +0.098, 95% CI 0.070-0.128), confirming that uncontrolled benchmarks over-credit agent-channel models. Missed-diagnosis detection was the universal weakness (best 0.467) and the one category where the human physicians outperformed every model. Raw detection ability does not guarantee a trustworthy score, and domestic and international models differ by deployment-relevant profile rather than by overall performance rank; both essential distinctions for performing clinical nuclear-medicine QC.
Hsu, C.-Y.; Liu, Q.; Shyr, Y.
Show abstract
As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.
Amiri, S.; Afshar, P.; Rohban, M. H.
Show abstract
Objectives. Radiomics pipelines extract hundreds of quantitative features that are widely known to be redundant, but the structure of this redundancy is usually treated as a per-dataset nuisance to be pruned away. We tested the alternative hypothesis that a substantial number of feature-feature correlations are universal: they persist across patients and across anatomically distinct structures because they reflect shared mathematical and image-statistical properties of how the image is summarised, rather than properties of the tissue being imaged. Materials and Methods. We re-analysed the publicly available Radiomics Atlas Dataset of normal Abdominal and Pelvic CT (RADAPT), restricting the analysis to the 526 non-contrast-enhanced examinations of the 531-subject atlas and to the 107 original (non-filtered) PyRadiomics features. The 53 segmented structures were grouped into four broad anatomical categories -- bones, muscles, vessels, and parenchymal organs. RADAPT is distributed as one Excel file per structure, with patients as rows and features as columns. Within each structure file we z-score-normalised every feature across patients, computed the absolute Spearman correlation matrix, and retained edges with |{rho}| [≥] {tau} for {tau} in {0.70, 0.80, 0.90}. We then intersected the edge sets across all structure files to obtain a "universal" correlation graph, in which an edge survives only if it exceeds the threshold in every structure (each estimated across the full patient sample). Stable feature communities were defined as the maximal cliques of this graph. Robustness to patient sampling was tested by repeating the entire pipeline on five independent random splits of each file into two patient halves (10 sub-cohorts per threshold), and the implementation was independently reproduced in R. Results. Despite the strictness of the global-intersection criterion, 34, 24, and 14 stable feature communities survived at {tau} = 0.70, 0.80, and 0.90 respectively, with the largest cliques containing six members at {tau} = 0.70 and {tau} = 0.80 and five members at {tau} = 0.90. The community structure was clearly interpretable: separate cliques captured (i) variance-like intensity dispersion, (ii) long-run / low-frequency (coarse) texture, (iii) high gray-level texture, (iv) low gray-level texture, (v) volume and surface shape, and (vi) local-homogeneity and energy/entropy duals. On random-half resampling the exact-match recovery rate of these communities was 81.5 %, 86.7 %, and 80.7 % across the three thresholds; departures from exact recovery were almost always a single boundary feature added or dropped, consistent with finite-sample fluctuation of near-threshold edges rather than structural instability. The R re-implementation reproduced the Python results exactly. Conclusion. A substantial portion of radiomics feature collinearity is universal across patients and tissues. We distinguish two layers within it: trivial near-algebraic duals that are universal by construction, and non-trivial cross-matrix-family communities that are the genuine empirical finding. Together they provide an interpretable, definition-grounded basis for aggressive dimensionality reduction, for retrospectively reconciling apparently different feature selections in the literature, and for moving radiomics pipelines toward organ-agnostic, more reproducible models. Clinical relevance statement. Selecting a single representative feature from each universal community shrinks the original-feature space by roughly an order of magnitude without sacrificing biologically distinct information. For example, the five variance-family members (first-order Variance, GLCM SumSquares, GLCM ClusterTendency, GLDM and GLRLM GrayLevelVariance) can be replaced by a single representative, removing redundant degrees of freedom that would otherwise inflate model variance; and labelling each retained feature by its community lets two studies that selected different variance-family names be recognised as having found the same signal, simplifying model development and improving cross-cohort generalisability in clinical CT workflows.
Jabbarpour, A.; Moulton, E.; Kaviani, S.; Zeng, W.; Ghassel, S.; Akbarian, R.; Couture, A.; Roy, A.; Liu, R.; Al-ali, Y.; Foufa, Y.; Hejji, N.; AlSulaiman, S.; Shirazi, Z.; Leung, E.; Klein, R.
Show abstract
Accurate interpretation of planar ventilation-perfusion (V/Q) scintigraphy, used for diagnosing pulmonary embolism (PE) based on PIOPED/EANM guidelines, requires objective assessment of mismatched V/Q defects. Manual delineation of V/Q defects is time-consuming, subject to interobserver variability, and rarely performed in practice, limiting standardized reporting and quantification of disease burden. To address these challenges, we evaluated four modern AI models for automated segmentation of vascular perfusion defects in planar V/Q scans and compared their performance to human annotators. We retrospectively identified 2,118 patients who underwent planar V/Q scans at The Ottawa Hospital (June 2019-February 2023). Six standard projections (ANT, POST, LAO, RAO, LPO, RPO) were included. Four 2D neural networks (U-Net, nnU-Net, Swin UNETR, and a Bottleneck Transformer U-Net [BTU-Net]) were trained on 1,313 patients (7,878 projections) and validated on 329 (1,974 projections) using physician-annotated defects. A hold-out test set of 46 high probability patients was used to evaluate segmentation quality, and defect detection accuracy using free-response receiver operating characteristic (FROC) analysis, where BTU-Net was the only model performing on par with human readers, showing robust sensitivity across the entire range of segmentation probabilities. At 1.5 false positives per projection rate (FPPR), BTU-Net outperformed other models with a sensitivity of 0.529 {+/-} 0.026, On a separate hold-out set of low likelihood of disease patients (n=430), the lowest FPPR was 0.08 {+/-} 0.01 for BTU-Net (P<0.0001). BTU-Net enables rapid, consistent, and accurate interpretation of planar V/Q scans. Such tools may enhance diagnostic efficiency, standardize reporting, and support non-expert readers in evaluating PE.
Blackman, B.; Fahey, N.; Dolan, S.; O'Reilly, M. K.; Cassidy, J. T.
Show abstract
Abstract Introduction: Proximal humerus fractures account for approximately 5-6% of all adult fractures and are primarily managed nonoperatively. Healing is conventionally monitored with radiographs, with radiopaque callus formation indicating healing. Visible radiographic callus appears weeks after biological union begins. Ultrasound provides a dynamic, radiation-free, and cost-effective method that can detect early callus formation before x-ray visibility. Although ultrasound has demonstrated utility for fracture healing in the clavicle and humeral shaft, its role in proximal humerus fractures remains unclear. Methods: This single-centre prospective study will be conducted in two phases. The pilot phase will measure inter-rater reliability for ultrasound detection of early callus formation at 2 and 4 weeks post-injury. Ten patients with proximal humerus fractures treated nonoperatively will undergo standardized anterior and lateral scans. Each patient will generate four saved images (short- and long-axis views), producing forty anonymized images independently reviewed by two raters. The prospective cohort phase will recruit approximately thirty additional patients. Results: Reliability will be quantified using Cohens kappa. A power calculation will be performed after pilot analysis. Results from the prospective cohort phase will help determine the association and predictive value of early ultrasound-detected bridging callus for radiographic and clinical union at three and six months. Patient reported outcome measures will be assessed using the Quick Disabilities of Arm, Shoulder and Hand (QuickDASH) questionnaire. Discussion: This study will develop and validate a standardized ultrasound protocol for assessing early fracture healing in proximal humerus fractures. By establishing both inter-rater reliability and predictive value, the findings may support ultrasound as a reproducible, radiation-free adjunct to conventional imaging and enable earlier identification of union status.
Li, Y.; Castelo, A.; Dennison, J. B.; Kettner, N. M.; Sieh, W.; Joseph, J. R.; Castillo, E.; Brock, K.; Weaver, O. O.; Wu, C.
Show abstract
Recent NCCN guideline highlighted AI-based mammographic risk prediction, but AI-based breast cancer detection remains questionable to translation. One barrier is current models often do not match routine clinical reasoning, which may add decision burden than benefits. In practice, radiologists compare current and prior mammograms while assessing breast density, bilateral symmetry, and lesion laterality. To align AI with this reasoning, we developed MuSTAF, a multi-task spatiotemporal attention fusion model for patient-level breast cancer classification from longitudinal full-field digital mammography. MuSTAF uses up to three recent mammograms, integrates temporal and cross-view information, refines suspicious-region features, and jointly predicts cancer status, breast density, and bilateral symmetry, with a separate laterality classifier for cancer-positive cases. In an internal case-control cohort (n = 351), MuSTAF achieved a cancer classification (AUC=0.84) exceeding all architecture-level baselines and published mammography AI models adapted to the same task (AUC [≤] 0.81). Simultaneously, it achieved AUCs of 0.83/0.80 for density/laterality assessments, and removing these auxiliary tasks reduced cancer detection performance. On the external CSAW-CC dataset (n = 8,723), model performance improved from 0.72 to 0.88 when restricting cancer cases to those with latest exams within 60 days before diagnosis, showing that temporally distant labels may shift detection evaluation toward risk prediction. Longitudinal analysis further showed that three recent exams outperformed five exams internally (AUC = 0.84 vs 0.80) and externally (0.72 vs 0.66), indicating recent imaging evidence mattered more than remote history. Overall, MuSTAF model improved longitudinal mammographic cancer classification while providing auxiliary outputs, and clarified temporal factors for applying AI to screening detection.
Bhattacharyya, K.
Show abstract
Designing transcutaneous skeletal muscle oxygenation (SmO2) sensors requires jointly optimizing source--detector geometry and wavelength selection while guaranteeing performance across populations that vary in subcutaneous fat thickness and skin pigmentation. We present a multi-fidelity Bayesian optimization (MFBO) framework that couples Monte Carlo light-transport simulations at two photon-count fidelities to a distributionally robust design objective. An autoregressive Gaussian-process surrogate learns the correlation between inexpensive low-photon-count and accurate high-photon-count simulations, and a cost-aware acquisition function decides both where and at what fidelity to sample. Robustness across the population is enforced with Conditional Value-at-Risk (CVaR) and entropic-risk (ERM) objectives that target worst-case subjects rather than the population average. On a five-layer forearm tissue model with anthropometric variability we find (i) a fidelity regime that is favorable for MFBO where the low-fidelity surrogate is rank-informative (Spearman {rho} = 0.84) but biased, at 100x lower cost; (ii) MFBO attains 23% higher robust sensitivity than a strong high-fidelity single-fidelity baseline at equal budget (p = 0.035), and avoids the optimistic bias that causes low-fidelity-only optimization to collapse when its designs are validated at high fidelity; (iii) CVaR/ERM objectives improve worst-case tail performance by {approx}23% relative to a mean objective without sacrificing average sensitivity; and (iv) discovered designs improve robust tail sensitivity by roughly 3--6x over commercial and heuristic optode layouts, with the largest gains in the high-fat and high-melanin subpopulations. The methodology bridges stochastic light-transport physics with sample-efficient machine-learning optimization and generalizes to cerebral oximetry, photodynamic therapy planning, and wearable physiological monitors.
Mandal, S.; Mendonca, P.; Gurushanth, K.; Thakur, H.; Birur, P.; Shetty, A.; Pal, D.
Show abstract
Background: The hyperplasia and dysplasia stage (pre-cancer) offers a viable opportunity to reduce the incidence and mortality of oral cancer through early prevention. Smartphone-based Artificial Intelligence (AI) enabled screening of potentially malignant oral lesions offers a scalable solution for this in resource-constrained settings. However, developing accurate and explainable AI segmentation models require high-quality, pixel-level annotated data. This process that is prohibitively expensive, time-consuming, and prone to inter-observer subjectivity among clinical experts. Methods: We designed and empirically validated a Deep Learning-driven Human-in-the-Loop (HITL) framework to pixel-annotate a dataset of 3026 clinical oral images. Using an iterative pseudo-labeling pipeline, we evaluated the model's learning dynamics and performance evolution across five training cycles. We conducted controlled experiments to quantify the networks tolerance to intermediate level of label noise (unreviewed pseudo-labels) to resolve clinical subjectivity using pixel-wise Cohen's Kappa and the STAPLE consensus algorithm. Results: Iterative self-training produced sustained improvements in lesion detection and spatial localization. However, controlled experiments revealed that including even a modest fraction ({approx}10%) of unreviewed pseudo-labels led to a three-to-four-fold increase in training convergence instability and induced a conservative prediction bias that negatively impacted model recall. When measured against multi-expert ground truth, the model's performance converged with the inter-rater reliability ceiling ({kappa} {approx} 0.65), indicating that its predictions fell within the envelope of human agreement. Conclusions: Our findings emphasize that a final expert-driven quality assurance step remains absolutely essential to mitigate training instability, confirmation bias, and clinically unacceptable drops in recall caused by label noise. Overall, this work provides a scalable, empirically validated blueprint for building domain-specific medical imaging datasets in low-resource global health settings, where the dual challenges of annotation cost and inter-observer variability are most acute.
Naidu, J. S.; Baskaradoss, V.
Show abstract
Background: Artificial intelligence (AI), including generative and foundation-based methods, has rapidly expanded within medical imaging research. However, the structure, citation impact, collaboration patterns, and thematic orientation of national research ecosystems remain incompletely characterised. Objectives: To evaluate global research trends in AI applied to medical imaging between 2017 and 2025, with detailed analysis of United Kingdom (UK)-affiliated output, citation performance, collaboration structure, funding landscape, and thematic evolution, with emphasis on generative and foundation-based methodologies. Materials and Methods: A bibliometric analysis of Scopus-indexed publications (2017-2025) was performed using a predefined search strategy targeting AI and medical imaging concepts, with emphasis on generative and foundation-based terms. Records were analysed globally and filtered for UK affiliation. Descriptive indicators including total publications (TP), total citations (TC), citations per paper (CPP), and year-on-year growth were calculated. Co-authorship and keyword co-occurrence networks were generated using VOSviewer (v1.6.19). Results: A total of 13,452 publications were identified globally (194,650 citations; global CPP 14.47), of which 889 (6.61%) were UK-affiliated. The UK ranked fourth by publication volume yet demonstrated higher citation efficiency (CPP 21.00) than several higher-volume countries. UK output increased approximately 18-fold between 2017 and 2025, with evidence of a citation-lag effect in recent years. Research activity was concentrated within a small number of institutions accounting for nearly half of national output, although citation impact varied independently of volume. Journal-dominant dissemination was associated with higher average citation impact compared with conference-centric models. Keyword analysis identified three principal thematic clusters: generative/deep learning methodologies, MRI- and diffusion-focused applications, and broader diagnostic imaging workflows. Highly cited publications were initially dominated by generative adversarial network-based reconstruction and synthesis, with recent rapid citation growth observed in diffusion and foundation-model architectures. Conclusion: UK-affiliated research represents a rapidly expanding and highly cited component of the global AI medical imaging literature, with increasing emphasis on generative, diffusion-based, and foundation-model approaches. These findings provide a reproducible bibliometric baseline for monitoring research activity, collaboration patterns, and potential translational priorities, while recognising that citation-based indicators do not directly measure clinical implementation, methodological quality, or real-world impact.